____ _ _ _ _
| _ \ ___ | |_ (_) _ __ ___ __| | (_) __ _
| |_) | / _ \ | __| | | | '_ \ / _ \ / _| | | | / _ |
| _ < | __/ | |_ | | | |_) | | __/ | (_| | | | | (_| |
|_| \_\ \___| \__| |_| | .__/ \___| \__,_| |_| \__,_|
|_|
- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b- `b
Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―Β―
Softmax-Funktion
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
top
In der Mathematik ist die sogenannte Softmax-Funktion oder normalisierte Exponentialfunktioncite-ref-2[1.1] eine Verallgemeinerung der logistischen Funktion, die einen K {\displaystyle K} -dimensionalen Vektor z {\displaystyle \mathbf {z} } mit reellen Komponenten in einen K {\displaystyle K} -dimensionalen Vektor Ο Ο ( z ) {\displaystyle \sigma (\mathbf {z} )} ebenfalls als Vektor reeller Komponenten in den Wertebereich ( 0 , 1 ) {\displaystyle (0,1)} transformiert, wobei sich die Komponenten zu 1 {\displaystyle 1} aufsummieren. Der Wert 1 {\displaystyle 1} kommt nur im Sonderfall K = 1 {\displaystyle K=1} vor. Die Funktion ist gegeben durch:
Ο Ο : R K β β ( 0 , 1 ) K {\displaystyle \sigma :\mathbb {R} ^{K}\to (0,1)^{K}}
Ο Ο ( z ) j = e z j β β k = 1 K e z k {\displaystyle \sigma (\mathbf {z} )_{j}={\frac {e^{z_{j}}}{\sum _{k=1}^{K}e^{z_{k}}}}} fΓΌr j = 1, β¦, K.
In der Wahrscheinlichkeitstheorie kann die Ausgabe der Softmax-Funktion genutzt werden, um eine kategoriale Verteilung β also eine Wahrscheinlichkeitsverteilung ΓΌber K {\displaystyle K} unterschiedliche mΓΆgliche Ereignisse β darzustellen. TatsΓ€chlich entspricht dies der gradient-log-Normalisierung der kategorialen Wahrscheinlichkeitsverteilung. Somit ist die Softmax-Funktion der Gradient der LogSumExp-Funktion.
Die Softmax-Funktion wird in verschiedenen Methoden der Multiklassen-Klassifikation verwendet, wie bspw. bei der multinomialen logistischen Regression (auch bekannt als Softmax-Regression)cite-ref-3[1.2]cite-ref-4[2], der multiklassen-bezogenen linearen Diskriminantenanalyse, bei naiven Bayes-Klassifikatoren und kΓΌnstlichen neuronalen Netzencite-ref-5[3]. Insbesondere in der multinomialen logistischen Regression sowie der linearen Diskriminantenanalyse entspricht die Eingabe der Funktion dem Ergebnis von K {\displaystyle K} distinkten linearen Funktionen, und die ermittelte Wahrscheinlichkeit fΓΌr die j {\displaystyle j} -te Klasse gegeben ein Stichprobenvektor x {\displaystyle x} und einem Gewichtsvektor w {\displaystyle w} entspricht:
P ( y = j β£ β£ x ) = e x T w j β β k = 1 K e x T w k {\displaystyle P(y=j\mid \mathbf {x} )={\frac {e^{\mathbf {x} ^{\mathsf {T}}\mathbf {w} _{j}}}{\sum _{k=1}^{K}e^{\mathbf {x} ^{\mathsf {T}}\mathbf {w} _{k}}}}}
Dies kann angesehen werden als Komposition von K {\displaystyle K} linearen Funktionen x β¦ β¦ x T w 1 , β¦ β¦ , x β¦ β¦ x T w K {\displaystyle \mathbf {x} \mapsto \mathbf {x} ^{\mathsf {T}}\mathbf {w} _{1},\ldots ,\mathbf {x} \mapsto \mathbf {x} ^{\mathsf {T}}\mathbf {w} _{K}} und der Softmax-Funktion (wobei x T w {\displaystyle \mathbf {x} ^{\mathsf {T}}\mathbf {w} } das innere Produkt von x {\displaystyle \mathbf {x} } und w {\displaystyle \mathbf {w} } bezeichnet). Die AusfΓΌhrung ist Γ€quivalent zur Anwendung eines linearen Operators definiert durch w {\displaystyle \mathbf {w} } bei Vektoren x {\displaystyle \mathbf {x} } , so dass dadurch die originale, mΓΆglicherweise hochdimensionale Eingabe in Vektoren im K {\displaystyle K} -dimensionalen Raum R K {\displaystyle \mathbb {R} ^{K}} transformiert wird.
Contents
β’ Alternativen
β’ Einzelnachweise
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Zusammenhang zur Logit-Funktion
Bei der binΓ€ren logistischen Regression benΓΆtigt man zur vollstΓ€ndigen Beschreibung lediglich die Wahrscheinlichkeit einer Klasse: P ( Y = 1 ) = 1 β β P ( Y = 0 ) {\displaystyle P(Y=1)=1-P(Y=0)} . FΓΌr zwei Klassen ist die Softmax-Funktion:
Ο Ο ( z ) j = e z j e z 1 + e z 2 {\displaystyle \sigma (\mathbf {z} )_{j}={\frac {e^{z_{j}}}{e^{z_{1}}+e^{z_{2}}}}} fΓΌr j = 1, 2 und Ο Ο ( z ) 2 = 1 β β Ο Ο ( z ) 1 {\displaystyle \sigma (\mathbf {z} )_{2}=1-\sigma (\mathbf {z} )_{1}} .
Da die z j {\displaystyle z_{j}} um eine beliebige Konstante verschoben werden kΓΆnnen ohne das Ergebnis zu Γ€ndern, gilt:
Ο Ο ( z ) 1 = e z 1 e z 1 + e z 2 = e z 1 e z 1 + e z 2 e β β z 2 e β β z 2 β β 1 = e z 1 e β β z 2 e z 1 e β β z 2 + 1 = e z ~ ~ e z ~ ~ + 1 = logit β β 1 β‘ β‘ ( z ~ ~ ) , {\displaystyle \sigma (\mathbf {z} )_{1}={\frac {e^{z_{1}}}{e^{z_{1}}+e^{z_{2}}}}={\frac {e^{z_{1}}}{e^{z_{1}}+e^{z_{2}}}}\underbrace {\frac {e^{-z_{2}}}{e^{-z_{2}}}} _{1}={\frac {e^{z_{1}}e^{-z_{2}}}{e^{z_{1}}e^{-z_{2}}+1}}={\frac {e^{\tilde {z}}}{e^{\tilde {z}}+1}}=\operatorname {logit} ^{-1}({\tilde {z}}),}
mit z ~ ~ = z 1 β β z 2 {\displaystyle {\tilde {z}}=z_{1}-z_{2}} und der Inversen der Logit-Funktion.
Alternativen
Softmax erzeugt Wahrscheinlichkeitsvorhersagen, welche ΓΌber ihrem TrΓ€ger dicht besetzt sind. Andere Funktionen wie sparsemax oder Ξ± Ξ± {\displaystyle \alpha } -entmax kΓΆnnen benutzt werden, wenn dΓΌnn besetzte Wahrscheinlichkeitsvorhersagen erzeugt werden sollencite-ref-6[4].
Einzelnachweise
1. Christopher M. Bishop: Pattern Recognition and Machine Learning. Springer, 2006 (englisch).
cite-note-42. β Computer Science Department: Unsupervised Feature Learning and Deep Learning Tutorial. Stanford University, abgerufen am 30. Januar 2019 (englisch).
cite-note-53. β Sophia Tamm: EinfΓΌhrung in neuronale Netze. In: Seminar Maschinelles Lernen - Dr. Zoran NikoliΔ. UniversitΓ€t KΓΆln, 30. Mai 2019, abgerufen am 24. Mai 2022.
cite-note-64. β Speeding Up Entmax, Maxat Tezekbayev, Vassilina Nikoulina, Matthias GallΓ©, Zhenisbek Assylbekov https://arxiv.org/abs/2111.06832v3